trusted data
Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
The growing importance of massive datasets with the advent of deep learning makes robustness to label noise a critical property for classifiers to have. Sources of label noise include automatic labeling for large datasets, non-expert labeling, and label corruption by data poisoning adversaries. In the latter case, corruptions may be arbitrarily bad, even so bad that a classifier predicts the wrong labels with high confidence. To protect against such sources of noise, we leverage the fact that a small set of clean labels is often easy to procure. We demonstrate that robustness to label noise up to severe strengths can be achieved by using a set of trusted data with clean labels, and propose a loss correction that utilizes trusted examples in a data-efficient manner to mitigate the effects of label noise on deep neural network classifiers. Across vision and natural language processing tasks, we experiment with various label noises at several strengths, and show that our method significantly outperforms existing methods.
Reviews: Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
Summary A method for learning a classifier robust to label noise is proposed. In contrast with most previous work, the authors take a simplifying but realistic assumption: there is of a small subset of the training set that has been sanitized (where labels are not noisy). Leveraging prior work, a noise transition matrix is first estimated and then used to correct the loss function for classification. Very extensive experiments support the validity of the method. Detailed comments This is a good contribution, although rather incremental with respect to Patrini et al. '17.
TrustNet: Learning from Trusted Data Against (A)symmetric Label Noise
Ghiassi, Amirmasoud, Younesian, Taraneh, Birke, Robert, Chen, Lydia Y.
Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. In this paper, we first derive analytical bound for any given noise patterns. Based on the insights, we design TrustNet that first adversely learns the pattern of noise corruption, being it both symmetric or asymmetric, from a small set of trusted data. Then, TrustNet is trained via a robust loss function, which weights the given labels against the inferred labels from the learned noise pattern. The weight is adjusted based on model uncertainty across training epochs. We evaluate TrustNet on synthetic label noise for CIFAR-10 and CIFAR-100, and real-world data with label noise, i.e., Clothing1M. We compare against state-of-the-art methods demonstrating the strong robustness of TrustNet under a diverse set of noise patterns.
Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
Hendrycks, Dan, Mazeika, Mantas, Wilson, Duncan, Gimpel, Kevin
The growing importance of massive datasets with the advent of deep learning makes robustness to label noise a critical property for classifiers to have. Sources of label noise include automatic labeling for large datasets, non-expert labeling, and label corruption by data poisoning adversaries. In the latter case, corruptions may be arbitrarily bad, even so bad that a classifier predicts the wrong labels with high confidence. To protect against such sources of noise, we leverage the fact that a small set of clean labels is often easy to procure. We demonstrate that robustness to label noise up to severe strengths can be achieved by using a set of trusted data with clean labels, and propose a loss correction that utilizes trusted examples in a data-efficient manner to mitigate the effects of label noise on deep neural network classifiers.
Trusted data will determine the future of baggage handling SITA
IATA sees RFID (radio frequency identification) as one of the keys to transforming the baggage handling process. SITA worked with IATA back in 2017 on a detailed business case, estimating that RFID could reduce the number of mishandled bags by an extra 25% and could potentially save the air transport industry $3 billion in baggage mishandling costs. Airlines and airports are now proactively working together to boost their baggage handling efforts as part of IATA's Resolution 753, which requires airlines to "maintain an accurate inventory of baggage by monitoring the acquisition and delivery of baggage". RFID tagging is now 99.98% accurate, according to IATA. Within the next four years most baggage systems will be RFID enabled, which is a huge improvement on barcodes alone.
Getting to Trusted Data via AI, Machine Learning, and Blockchain
Establishing trust in data is an essential requirement for businesses and entities for whom credible, reliable information is the lifeblood. As enterprises seek to manage data as an asset, it becomes increasingly vital that data sources are trusted and verifiable. I wrote a few weeks ago about the MIT initiative to establish a framework for trusted data, and the resulting position paper, "Towards an Internet of Trusted Data: A New Framework for Identity and Data Sharing". The authors highlight the criticality and need for "trustworthy, auditable data provenance" where "systems must automatically track every change that is made to data, so it is auditable and completely trustworthy". One of the key recommendations of the study was to improve the process and quality of data sharing.
Getting to Trusted Data via AI, Machine Learning, and Blockchain
Establishing trust in data is an essential requirement for businesses and entities for whom credible, reliable information is the lifeblood. As enterprises seek to manage data as an asset, it becomes increasingly vital that data sources are trusted and verifiable. I wrote a few weeks ago about the MIT initiative to establish a framework for trusted data, and the resulting position paper, "Towards an Internet of Trusted Data: A New Framework for Identity and Data Sharing". The authors highlight the criticality and need for "trustworthy, auditable data provenance" where "systems must automatically track every change that is made to data, so it is auditable and completely trustworthy". One of the key recommendations of the study was to improve the process and quality of data sharing.
Getting to Trusted Data via AI, Machine Learning, and Blockchain
Establishing trust in data is an essential requirement for businesses and entities for whom credible, reliable information is the lifeblood. As enterprises seek to manage data as an asset, it becomes increasingly vital that data sources are trusted and verifiable. I wrote a few weeks ago about the MIT initiative to establish a framework for trusted data, and the resulting position paper, "Towards an Internet of Trusted Data: A New Framework for Identity and Data Sharing". The authors highlight the criticality and need for "trustworthy, auditable data provenance" where "systems must automatically track every change that is made to data, so it is auditable and completely trustworthy". One of the key recommendations of the study was to improve the process and quality of data sharing.